Papers with distributional model
Can a Gorilla Ride a Camel? Learning Semantic Plausibility from Text (D19-60)
Copied to clipboard
| Challenge: | Existing work on modeling semantic plausibility has focused on physical plausability but distributional methods fail when tested in supervised settings. |
| Approach: | They propose to use large pretrained language models to model plausibility in supervised settings by extracting attested events from a large corpus and injecting explicit commonsense knowledge into a distributional model. |
| Outcome: | The proposed model is effective in modeling plausibility in a supervised setting. |
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)
Copied to clipboard
| Challenge: | DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing . |
| Approach: | They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus . |
| Outcome: | The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset. |
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)
Copied to clipboard
| Challenge: | A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language. |
| Approach: | They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far. |
| Outcome: | The proposed model has the largest vocabulary and the best evaluation scores published so far. |
When Hearst Is not Enough: Improving Hypernymy Detection from Corpus with Distributional Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | a taxonomy is a semantic hierarchy of words or concepts organized w.r.t. their hypernymy relationships. |
| Approach: | They propose a framework for hypernymy detection using large textual corpora . they quantify the non-negligible existence of specific sparsity cases . |
| Outcome: | The proposed framework quantifies the non-negligible existence of specific sparsity cases on several benchmark datasets. |
A Multi-word Expression Dataset for Swedish (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing data on compositionality of multi-word expressions is limited and only available for high resource languages. |
| Approach: | They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression . |
| Outcome: | The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality. |